Phase 3: Machine Learning Lesson 6 of 6

Evaluation Metrics:
How Good is Your Model, Really?

Accuracy is just the beginning. Different problems require different yardsticks, and picking the wrong metric can lead you to build a model that looks impressive on paper but fails completely in practice.

You will learn
Why accuracy alone can deceive you
Precision, recall, and F1 score in depth
What the ROC curve and AUC-ROC measure
Regression evaluation: MAE, RMSE, and R-squared
How to choose the right metric for any problem

The problem with accuracy

Accuracy is defined as the fraction of correct predictions out of all predictions. Simple, intuitive, easy to explain. It is also deeply misleading in many of the situations where you most need a reliable metric.

Consider a fraud detection system. In a typical dataset, maybe 0.5% of transactions are fraudulent. A model that predicts "not fraud" for every single transaction achieves 99.5% accuracy. It catches zero fraudulent transactions. Nobody would use it. Yet accuracy says it is excellent.

Analogy

Imagine rating a smoke alarm purely by how often it is silent. A smoke alarm that never beeps achieves a perfect silence rate. But it also fails to do its one job: alerting you when there is a fire. Accuracy on imbalanced datasets is the smoke alarm version of model evaluation. It measures the wrong thing entirely.

The moment your classes are imbalanced, which they almost always are in real fraud, medical, and security applications, you need better tools. This lesson gives you those tools.

Classification metrics that actually matter

Everything you need starts from the confusion matrix. You met it in Lesson 3.2. Now let's build the key metrics from its four cells: True Positives (TP), True Negatives (TN), False Positives (FP), and False Negatives (FN).

Precision
Exactness
TP / (TP + FP). Of every time the model said "positive," how often was it correct? High precision means few false alarms.
Recall
Coverage
TP / (TP + FN). Of every actual positive in the data, how many did the model catch? High recall means few missed cases.
F1
Harmonic mean
2 × (Precision × Recall) / (Precision + Recall). Single number balancing both. Penalises large gaps between precision and recall.
AUC
Area under ROC
Threshold-independent measure of a classifier's ability to separate classes. 0.5 = random guessing; 1.0 = perfect classifier.

The precision-recall tradeoff

Precision and recall are locked in a tug-of-war. If you make your model more aggressive about flagging positives, recall goes up (you catch more real cases) but precision goes down (you also raise more false alarms). If you make it more conservative, precision goes up but recall falls.

Which end of that tradeoff you prefer depends entirely on the cost of each type of error in your specific context.

Context Worse error Optimise for Why
Cancer screening False Negative (missed cancer) High Recall Missing a real cancer case has far greater consequences than unnecessary follow-up tests.
Spam filtering False Positive (real email flagged as spam) High Precision Users accept some spam in their inbox more readily than they accept missing important emails.
Fraud detection False Negative (missed fraud) High Recall Each missed fraud costs money. False positives mean some legitimate transactions get reviewed manually.
Content recommendation False Positive (irrelevant recommendation) High Precision Showing bad recommendations erodes user trust quickly. It is better to show fewer but more relevant items.

When you cannot clearly prioritise one over the other, use F1. When your problem has a heavy class imbalance, consider also reporting the balanced accuracy or the Matthews Correlation Coefficient, which account for the imbalance directly.

The ROC curve and AUC-ROC

Most classifiers do not just output a class label. They output a probability. "This email has a 91% chance of being spam." You then apply a threshold: anything above 0.5 gets labelled spam. But that threshold is a choice. Different thresholds give different precision-recall tradeoffs.

The ROC (Receiver Operating Characteristic) curve visualises the full range of this tradeoff across every possible threshold. It plots True Positive Rate (recall) against False Positive Rate (1 minus specificity) as the threshold sweeps from 0 to 1. A perfect classifier hugs the top-left corner. A random classifier follows the diagonal.

ROC curves: comparing three classifiers
False Positive Rate True Positive Rate (Recall) 0.0 0.3 0.7 1.0 0.0 0.3 0.7 1.0 Random (AUC=0.5) Excellent (AUC = 0.97) Good (AUC = 0.85) Weak (AUC = 0.65)

The further the curve bends toward the top-left corner, the better the classifier at separating classes regardless of threshold. AUC (Area Under the Curve) summarises this into a single number you can use to compare models directly.

AUC-ROC is especially useful because it is threshold-independent. You can compare two models on their AUC without committing to a particular decision threshold upfront. Once you have selected the best model by AUC, you then choose the threshold based on your precision-recall priorities for the deployment context.

AUC at a glance

AUC = 1.0: Perfect. The model correctly separates every positive from every negative. AUC = 0.9: Excellent. The model would correctly rank a random positive above a random negative 90% of the time. AUC = 0.7: Acceptable for some tasks; investigate further. AUC = 0.5: Random guessing. Your model has learned nothing useful.

Regression evaluation metrics

Regression problems have their own set of metrics, because error in continuous outputs cannot be measured as right or wrong. Every prediction has a magnitude of error, and different metrics weight those magnitudes differently.

MAE
Mean Absolute Error
Average size of errors in the same units as the target. Intuitive and robust to outliers. "The model is off by an average of $4,200."
RMSE
Root Mean Squared Error
Square root of the average squared error. Penalises large errors more than MAE. Also in the same units as the target but more sensitive to outliers.
R²
R-Squared
Proportion of variance explained by the model. 1.0 is perfect; 0.0 means the model does no better than simply predicting the mean every time.

When you report regression results to a non-technical audience, MAE is your friend. "Our model predicts delivery time with an average error of 8 minutes" is something any stakeholder can understand immediately. RMSE is harder to interpret but better for comparing models with each other because it penalises bad predictions heavily. R-squared gives the most complete picture but can be misleading on non-linear data.

Putting it all together in code

Here is a single Python snippet that generates a full evaluation report for a classification model: accuracy, precision, recall, F1, and the ROC-AUC score. Copy this pattern and use it on every classification project.

Python full_evaluation.py
from sklearn.datasets import load_breast_cancer
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score,
    f1_score, roc_auc_score, classification_report
)

# Load data and train a model
X, y = load_breast_cancer(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42, stratify=y
)

model = RandomForestClassifier(n_estimators=100, random_state=42)
model.fit(X_train, y_train)

# Hard predictions (class labels)
preds = model.predict(X_test)

# Soft predictions (probabilities) for AUC
probs = model.predict_proba(X_test)[:, 1]

# Print full report
print("=== Classification Report ===")
print(f"Accuracy:  {accuracy_score(y_test, preds):.4f}")
print(f"Precision: {precision_score(y_test, preds):.4f}")
print(f"Recall:    {recall_score(y_test, preds):.4f}")
print(f"F1 Score:  {f1_score(y_test, preds):.4f}")
print(f"AUC-ROC:   {roc_auc_score(y_test, probs):.4f}")
print()
print(classification_report(y_test, preds, target_names=['Malignant', 'Benign']))
Output
=== Classification Report ===
Accuracy: 0.9649
Precision: 0.9730
Recall: 0.9730
F1 Score: 0.9730
AUC-ROC: 0.9955

                precision   recall  f1-score   support
Malignant       0.95      0.95      0.95      42
   Benign       0.97      0.97      0.97      72

The AUC-ROC of 0.9955 is exceptional. Even if we lower the decision threshold to catch more malignant cases at the cost of some false alarms, this model will almost certainly outperform any reasonable baseline. The F1 scores for both classes are above 0.95, confirming the model is not just winning because one class dominates the data.

🎓
Phase complete

You have finished Phase 3: Machine Learning

You started this phase not knowing what a model actually does. You now understand the learning loop, classification, regression, clustering, how overfitting works and how to fight it, and how to measure whether a model is genuinely useful. That is a solid foundation. Phase 4 takes everything you have built and extends it into neural networks and deep learning.

Final Phase 3 activity

Build and evaluate a production-ready pipeline

You are going to build a complete ML pipeline on a real imbalanced dataset, compare three algorithms using cross-validated AUC scores, tune the best one, set a threshold based on your precision-recall priorities, and produce a final report you could actually present to a stakeholder.

01 Open the Lesson 3.6 Colab notebook. Load the credit card fraud dataset and check the class balance. What percentage of transactions are fraud?
02 Train three models: LogisticRegression, RandomForestClassifier, and GradientBoostingClassifier. Compare their 5-fold cross-validated AUC scores. Which wins?
03 Plot the ROC curve for the best model. Also plot the Precision-Recall curve. Given that missing a fraud is more costly than a false alert, should you raise or lower the threshold?
04 Set a threshold that achieves at least 90% recall. What happens to precision? What is the F1 score at this threshold?
05 Write a 5-sentence plain-English summary of your model's performance, as if presenting to a bank's fraud team. Avoid jargon. Use your MAE/recall/precision numbers as evidence.
Your Notes
Studying independently? Write your thoughts or answers below. Notes save automatically to your browser.
Practice Notebook
Run this lesson's code live in Google Colab
All examples + challenge exercises · Free GPU included · No setup required
Open In Colab
Progress
Done with this lesson?
Mark it complete to track your progress.